Papers by Kristine Ann M. Carandang

4 papers
PathBuilder: A Quality-Controlled LLM System for Personalized Learning Pathways (2026.acl-demo)

Copied to clipboard

Challenge: Large language models (LLMs) enable scalable content generation for personalized learning, but reliability and pedagogical alignment remain open challenges.
Approach: They propose a web-based system that integrates expert-validated assessment, retrieval-augmented generation (RAG), and an LLM-as-a-Judge validation loop within a closed instructional pipeline.
Outcome: The proposed system achieves a gain of 37.9 percentage points and a large effect size in a real-world deployment with 179 registered users.
When Models Hesitate: Answer Instability as a Label-Free Uncertainty Signal for LLMs (2026.acl-srw)

Copied to clipboard

Challenge: Existing approaches to uncertainty estimation typically require access to internals, additional supervision, or computationally intensive pipelines.
Approach: They propose to use a label-free uncertainty signal to predict the variability of a model's final answer across repeated stochastic generations of the same prompt to achieve performance competitive with semantic entropy.
Outcome: The proposed method achieves performance competitive with semantic entropy while requiring no similarity model.
Foundations of PEERS: Assessing LLM Role Performance in Educational Simulations (2025.acl-srw)

Copied to clipboard

Challenge: In education, peer instruction is widely recognized as an effective active learning strategy, but evaluations of PI are limited by logistical constraints and variability in classroom settings.
Approach: They propose a simulation framework that integrates Agent-Based Modeling, Large Language Models, and Bayesian Knowledge Tracing to emulate student learning dynamics.
Outcome: The proposed framework integrates Agent-Based Modeling, Large Language Models, and Bayesian Knowledge Tracing to emulate student learning dynamics in real classrooms.
Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation (2025.acl-industry)

Copied to clipboard

Challenge: Clinical note generation (CNG) tools are being developed to address extended working hours and healthcare provider fatigue.
Approach: They evaluate the reliability of 12 open-weight and proprietary LLMs from Anthropic, Meta, Mistral, and OpenAI in CNG in terms of their ability to generate notes that are string equivalent (consistency rate), have the same meaning (semantic consistency) and are correct (symbol similarity)
Outcome: The results show that the LLMs generated notes that are string equivalent (consistency rate), have the same meaning (semantic consistency) and are correct (symbol similarity) overall, Meta’s Llama 70B was the most reliable, followed by Mistral’s Small model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations